> grep -r "local LLM"▋
8 entries
local LLM: 11 posts on this blog. The newest is “llama.cpp vs Ollama — which one should you run?” (6 October 2026). Auto-tended from the posts below — every sentence cites its source.
auto-tended from 11 posts · every sentence cites its source · mortal on purpose
2026-10-06
llama.cpp vs Ollama — which one should you run?
The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.
2026-10-02
GGUF VRAM Calculator: Check Before You Download
Model size, quant and context in — weights, KV cache and a per-card verdict out. The GGUF VRAM calculator answers will-it-fit before the download starts.
2026-09-23
Can vLLM Run GGUF? Yes — on GPU Only
Can vLLM run GGUF? Yes — via the official plugin, on GPU only. The serve syntax, the tokenizer trap, the hardware limits, and when llama.cpp still wins.
2026-09-13
GGUF Quantization Levels: Q4_K_M vs Q8_0, Size and VRAM
GGUF quantization levels compared: Q4_K_M vs Q8_0 on size, VRAM and quality, plus the one rule — default Q4_K_M, step up only for code and math.
2026-09-10
llama.cpp vs Ollama: Which Should You Run in 2026?
llama.cpp vs Ollama: Ollama wraps llama.cpp for convenience, raw llama.cpp wins on speed and control. Benchmarks, GPU offload flags, and when each wins.
2026-09-07
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
How to run GGUF models locally: one-line Ollama pulls, llama.cpp straight off a Hugging Face URL, and how to pick the right quant for your VRAM.
2026-09-04
Best Local LLM for Coding: 8GB to 24GB VRAM Picks
The best local LLM for coding by VRAM bracket: Qwen3 Coder vs DeepSeek at 8, 12, 16 and 24 GB, the right quant per card, plus a runnable Ollama setup.
2026-09-01
Ollama vs LM Studio: Which Local LLM Tool Should You Use?
Ollama vs LM Studio compared for real dev work: install, GPU use, speed and API serving — with runnable commands so you can pick the right tool today.